Papers with language modeling
Copied to clipboard
| Challenge: | GluonNLP is a powerful new toolkit that automates the most laborious aspects of deep learning for NLP. |
| Approach: | This hands-on tutorial demonstrates how to scale unsupervised pre-training techniques with Apache MXNet and GluonNLP. |
| Outcome: | This hands-on tutorial examines the challenges of scaling these models and algorithms effectively with Apache MXNet and GluonNLP. |
Copied to clipboard
| Challenge: | Latent structure models are a powerful tool for compositional data modeling and pipelines. |
| Approach: | This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations . |
| Outcome: | This tutorial will cover recent advances in discrete latent structure models . it will discuss their motivation, potential, and limitations . |
Copied to clipboard
| Challenge: | Pretraining by language modeling has become popular but we have yet to understand what language models learn during that process. |
| Approach: | They propose diagnostics that ask questions about information used by language models for generating predictions in context. |
| Outcome: | The proposed diagnostics can be used to study the popular BERT model . they show that the model can distinguish good from bad completions, but struggles with inference and role-based event prediction. |
Copied to clipboard
| Challenge: | Recent studies have determined that the learned token embeddings of large-scale neural language models are degenerated to be anisotropic with a narrow-cone shape. |
| Approach: | They propose a method to degenerate the learning gradient for rare token embeddings by gating the specific part of the gradient for all tokens during training stage. |
| Outcome: | The proposed method improves the performance of the models but lacks the training dynamics needed to solve the representation degeneration problem. |
Copied to clipboard
| Challenge: | Named entities pose a unique challenge to traditional methods of language modeling. |
| Approach: | They propose a Hierarchically Disentangled Model for named entities in cooking recipes using a dataset from several publicly available online sources. |
| Outcome: | The proposed model is based on 158,473 cooking recipes from public sources. |
Copied to clipboard
| Challenge: | Existing studies show that syntactic information is useful for a wide variety of NLP tasks. |
| Approach: | They propose to use word-level representations to learn internal representations that capture soft hierarchical notions of syntax from highly varied supervision. |
| Outcome: | The proposed model encodes significant amounts of syntax even without explicit supervision. |
Copied to clipboard
| Challenge: | Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations . |
| Approach: | They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest. |
| Outcome: | The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 . |
Copied to clipboard
| Challenge: | Neural codec language models (or codec LMs) are emerging as a powerful framework for text-to-speech (TTS) despite the close interdependence of codecs and LM, research on codec and lms has largely remained siloed. |
| Approach: | They propose a frame-wise codec encoder that improves both LM log-likelihood and TTS metrics . they also propose LM codebook level dropout to efficiently navigate a portion of codec-LM design space . |
| Outcome: | The proposed codec-LM co-design improves intelligibility, audio quality and speaker control compared to a siloed baseline. |
Copied to clipboard
| Challenge: | Scaling laws in language modeling quantify training loss as a function of dataset size and model parameters, but neglect the critical role of data quality in model generalization. |
| Approach: | They propose to use effective training tokens as a combination of text diversity and syntheticity as measured by a teacher model to calculate scaling laws. |
| Outcome: | The proposed term effective training tokens is a combination of two readily-computed indicators of text diversity and syntheticity as measured by a teacher model. |
Copied to clipboard
| Challenge: | OpenNMT is a community-built toolkit written in multiple languages with an emphasis on extensibility. |
| Approach: | They propose to use PyTorch to train custom sequence models for translation, summarization, language modeling, and other tasks. |
| Outcome: | The proposed toolkit is fast, extensible, and useful for both research and production. |
Copied to clipboard
| Challenge: | XLM-R models encode language-sensitive information in each language, allowing them to extract features for downstream tasks and cross-lingual transfer learning. |
| Approach: | They evaluate how multilingual language models maintain a shared multilingual representation space while still encoding language-sensitive information in each language. |
| Outcome: | The proposed model can extract features for downstream tasks and cross-lingual transfer learning. |
Copied to clipboard
| Challenge: | RENarGen generates closed narratives by ensuring the first and last sentences are related and then infilling the middle sentences. |
| Approach: | They propose a novel novel novel that generates closed narratives by ensuring the first and last sentences are related and then infilling the middle sentences. |
| Outcome: | The proposed paradigm generates closed narratives by ensuring the first and last sentences are related and then infilling the middle sentences. |
Copied to clipboard
| Challenge: | Variational autoencoders (VAEs) have been widely applied in text generation tasks, but they suffer from insufficient representation capacity and poor controllability. |
| Approach: | They propose a data-driven prior that has expressivity and controllability. |
| Outcome: | The proposed prior enjoys expressivity and controllability and can be used in language modeling and controlled text generation. |
Copied to clipboard
| Challenge: | Existing methods for length extrapolation are tailored for natural language modeling, a task known to have strong recency bias. |
| Approach: | They propose two attention alignment strategies to improve T5's long-context utilization capability without fine-tuning. |
| Outcome: | The proposed methods improve the long-context utilization capability of T5 on language modeling, retrieval, multi-document question answering, and code completion tasks without any fine-tuning. |
Copied to clipboard
| Challenge: | In contrast, adversarial attacks can cause model errors by modifying inputs, such as the universal triggers attack. |
| Approach: | They propose a data poisoning attack that allows an adversary to control model predictions whenever a desired trigger phrase is present in the input. |
| Outcome: | The proposed attack can cause model errors by modifying inputs, but it can also cause extra human annotation. |
Copied to clipboard
| Challenge: | Existing explanation methods conflate evidence for various features to predict a token . existing explanation methods are less interpretable for human understanding . |
| Approach: | They propose to explain language models contrastively by looking for salient input tokens that explain why the model predicted one token instead of another. |
| Outcome: | The proposed explanations are better than non-contrastive explanations for language models . they show that contrastive explanations improve simulability for human observers . |
Copied to clipboard
| Challenge: | Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model. |
| Approach: | They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model. |
| Outcome: | The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling. |
Copied to clipboard
| Challenge: | Autoregressive generative models are often criticized for using ground-truth contexts at training time but generated ones at test time. |
| Approach: | They propose that generalization is the underlying property to address and propose unconditional generation as its fundamental benchmark. |
| Outcome: | The proposed model is generalized and can handle true and generated contexts. |
Copied to clipboard
| Challenge: | Neural-based approaches to natural language generation are data-hungry and difficult to adopt in real-world applications. |
| Approach: | They propose a task of few-shot natural language generation from structured data or knowledge to generate coherent sentences from input data and language modeling to compose coherent sentences. |
| Outcome: | The proposed approach outperforms the strongest baseline approach by over 8.0 BLEU points improvement. |
Copied to clipboard
| Challenge: | Existing languages have syntactic representations of code to improve code intelligence, but they are difficult to learn from code. |
| Approach: | They propose to embed dynamic information of programs revealed by their test cases into feature representations of code as complements. |
| Outcome: | The proposed method yields 6%/19% mAP improvements over its masked language modeling counterparts. |
Copied to clipboard
| Challenge: | Variational autoencoders (VAEs) with an auto-regressive decoder have been applied for many natural language processing tasks. |
| Approach: | They propose a cyclical annealing schedule which repeats the process of increasing multiple times to learn more meaningful latent codes progressively by leveraging previous learning cycles as warm re-restart. |
| Outcome: | The proposed method improves on a broad range of NLP tasks, including language modeling, dialog response generation and semi-supervised text classification. |
Copied to clipboard
| Challenge: | Low-resource languages (LRLs) face significant challenges in natural language processing due to limited data. |
| Approach: | They evaluate adapter-based methods for adapting mLMs to low-resource languages . they use unstructured text and structured knowledge from ConceptNet to evaluate adapters . |
| Outcome: | The proposed methods outperform large language models and LLaMA-3 and deepSeek-R1 models on low training data. |
Copied to clipboard
| Challenge: | GPT is an auto-regressive Transformer-based pre-trained language model . but its huge size can be prohibitive for deploying on low capacity devices . |
| Approach: | They use a Kronecker decomposition technique to compress GPT models . they use ILKD to refine the model on downstream tasks . |
| Outcome: | The proposed model outperforms the existing DistilGPT2 model on language modeling and general language understanding evaluation benchmark tasks. |
Copied to clipboard
| Challenge: | Efficient transformer variants with linear time complexity have been developed to mitigate the quadratic computational overhead of the vanilla transformer. |
| Approach: | They propose a linear time complexity transformer variant that reduces the quadratic computational overhead of the vanilla transformer by using a recurrent-style incremental computation similar to kernel-based transformers. |
| Outcome: | The proposed method reduces the performance gap while achieving the same efficiency even with short generation. |
Copied to clipboard
| Challenge: | Sentence pair modeling is critical for many NLP tasks, such as paraphrase identification and semantic textual similarity. |
| Approach: | They propose to use subwords to represent sentences without pretrained word embeddings . they find that subword models can achieve new state-of-the-art results without pretraining . |
| Outcome: | The proposed models can achieve state-of-the-art results on two social media datasets and competitive results on news data for paraphrase identification. |
Copied to clipboard
| Challenge: | Empirical experiments show that our model learns latent distributions that respect latent space geometry and is able to generate sentences that are more diverse. |
| Approach: | They propose a Variational Wasserstein Autoencoder with Riemannian Normalizing Flow to solve this problem by transforming a latent variable into a space that respects the geometric characteristics of input space. |
| Outcome: | Empirical results show that the proposed model avoids KLvanishing and has better performance in language modeling, likelihood approximation, and text generation tasks. |
Copied to clipboard
| Challenge: | Retrieval-augmented language models are a promising alternative to standard pretraining, but little attention has been put into understanding what this type of training scheme does to the underlying language model when analyzed as a standalone -separated from the overall retrieval pipeline. |
| Approach: | They propose an ‘ideal retrieval’ methodology to study these models in a fully controllable setting and propose a retrieval augmentation methodology to examine their effects. |
| Outcome: | The proposed model saves substantially less world knowledge in their weights, but is worse at comprehending global context. |
Copied to clipboard
| Challenge: | Currently, on-device keyboards have limited memory and response time for word prediction . a proposed on-device neural language model based word prediction method is available for mobile devices . |
| Approach: | They propose an on-device neural language model based word prediction method that optimizes run-time memory and provides a real-time prediction environment. |
| Outcome: | The proposed model outperforms existing methods for word prediction in keystroke savings and word prediction rate and has been commercialized. |
Copied to clipboard
| Challenge: | a key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language. |
| Approach: | They propose to use a full-vocabulary setup to test the performance of language modeling (LM) on 50 typologically diverse languages. |
| Outcome: | The proposed language modeling task is based on a full vocabulary setup focused on word-level prediction on 50 typologically diverse languages. |
Copied to clipboard
| Challenge: | Existing approaches to de-bias pre-trained large language models focus on changes to training regime, but this is not feasible. |
| Approach: | They propose to de-bias a pre-trained model by fine-tuning it on only 10 examples . they show that the technique performs better than competitive baselines . |
| Outcome: | The proposed method performs better than competitive state-of-the-art baselines with minimal loss in language modeling ability. |
Copied to clipboard
| Challenge: | Autoregressive Transformer language models do not require explicit positional encodings (PEs) this is because a cascade of (permutation invariant) set processors collectively exhibit sequence-sensitive behavior in the autoregressively setting. |
| Approach: | They propose to explain why autoregressive Transformers require explicit positional encodings (PEs) this property has been known since early efforts adopting the Transformer for language modeling . |
| Outcome: | The proposed model can distinguish sequences with permuted tokens without the need for explicit PEs. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Language Modeling (RALM) is a popular approach for large language models. |
| Approach: | They propose a modular RALM that integrates large language models with documents from an external corpus to improve inference efficiency. |
| Outcome: | The proposed method improves inference efficiency with appending context pattern while maintaining decent performance after fine-tuning by Low-Rank Adaption. |
Copied to clipboard
| Challenge: | Adversarial attacks against Language models (LMs) are a significant concern. |
| Approach: | They propose an approach to automatically learn a policy to generate challenging examples that improve the model’s performance. |
| Outcome: | The proposed approach outperforms baselines and exhibits generalizability across classifiers and datasets. |
Copied to clipboard
| Challenge: | Recent studies have shown that sharing key-value (KV) cache across layers is effective in efficient inference of large language models. |
| Approach: | They propose a unified framework that covers several recent methods and their novel variants to investigate cross-layer KV sharing. |
| Outcome: | The proposed framework achieves higher throughput and better performance when reducing the size of the key-value cache by 2 while maintaining competitive performance. |
Copied to clipboard
| Challenge: | Recent studies show that encoding more syntactic information does not lead to better performance. |
| Approach: | They propose a method to optimize pareto-optimal models by formalizing it as a multi-objective optimization problem. |
| Outcome: | The proposed method is better than a baseline method on two NLP tasks. |
Copied to clipboard
| Challenge: | Existing models that learn embeddings only in Euclidean vector space do not account for such structural property of language. |
| Approach: | They propose a Poincare Variational Autoencoder to capture latent hierarchies in hyperbolic space . they propose enabling adversarial learning procedures to empower robust model training . |
| Outcome: | The proposed model outperforms existing models in a hyperbolic latent space . it captures latent language hierarchies in hyperbolical space and is robust to training . |
Copied to clipboard
| Challenge: | a new study examines the impact of NLP research published in top-tier conferences from 1979 to 2024 . language modeling has the widest internal and external influence, while linguistic foundations have lower impacts . |
| Approach: | They analyze citations from research articles and external sources to determine how NLP topics are consumed internally and externally. |
| Outcome: | The findings show that language modeling has the widest internal and external influence . ethics, bias, and fairness show significant attention in policy documents with fewer academic citations . |
Copied to clipboard
| Challenge: | Efficient fine-tuning of large language models requires non-trivial efforts to implement these methods on different models. |
| Approach: | They propose a framework that democratizes the fine-tuning of large language models by integrating a suite of efficient training methods into one framework. |
| Outcome: | The proposed framework is able to scale to 100+ LLMs without coding and receives over 25,000 stars and 3,000 forks. |
Copied to clipboard
| Challenge: | Current language models achieve low perplexity but their resulting generations still suffer from toxic responses, repetitiveness, and contradictions. |
| Approach: | They propose a new language model architecture that uses a language modeling and a classification head for each output token. |
| Outcome: | The proposed model outperforms existing model guiding approaches in terms of accuracy and efficiency. |
Copied to clipboard
| Challenge: | a recent attempt at language modeling with predicted semantic structure failed to establish empirical lower bounds on what could have made the attempt successful. |
| Approach: | They propose a concise binary vector representation of semantic structure at the lexical level and evaluate how good an incremental tagger needs to be to achieve better-than-baseline performance. |
| Outcome: | The proposed model can achieve better-than-baseline performance without losing its main advantages and lower bounds on prediction quality can't be established via a single score alone. |
Copied to clipboard
| Challenge: | Existing models trained on poor quality data have shown strong performance in language modeling and some downstream benchmarks. |
| Approach: | They evaluate kNN-LMs on a diverse set of tasks and evaluate their performance. |
| Outcome: | The proposed extension could improve on a variety of tasks, but it fails to perform on reasoning tasks that require integrating multiple pieces of information. |
Copied to clipboard
| Challenge: | Existing frameworks for natural language processing ignore interactions among different heads, which wastes the capacity of the model. |
| Approach: | They propose a model which explicitly models interactions between attention heads through a hierarchical variational distribution. |
| Outcome: | The proposed model outperforms the baseline model on Wikitext-103 and WMT14 EN-DE on language modeling and translation tasks. |
Copied to clipboard
| Challenge: | Language modeling is a fundamental task in natural language processing, applications include machine translation, image captioning and speech recognition. |
| Approach: | They propose a cosine regularization method to solve the representation degeneration problem by analyzing the limitations of the proposed method and then propose an alternative regularization technique to tackle the problem. |
| Outcome: | The proposed method is effective in language modeling and image captioning. |
Copied to clipboard
| Challenge: | Recent research shows that retrieval-augmented models with shorter contexts (4K tokens) can match the performance of models with longer contexts (16K/32K token) |
| Approach: | They introduce an approach to extend the effective context size of large language models by using an external vector cache to store past states. |
| Outcome: | The proposed method improves on models trained from scratch and pre-trained models. |
Copied to clipboard
| Challenge: | General-purpose models lack depth for expert-level tasks because of limited domain-specific information. |
| Approach: | They propose a method for curating domain-specific datasets from noisy web sources to improve model performance. |
| Outcome: | The proposed model outperforms the baseline model on the astronomy benchmark and on the AstroBench. |
Copied to clipboard
| Challenge: | Existing language modeling models treat text sequences as if they were created independently. |
| Approach: | They propose a hierarchical extension to the language modeling problem whereby a human-level exists to connect sequences of documents and capture the notion that human language is moderated by changing human states. |
| Outcome: | The proposed model outperforms the current state-of-the-art in terms of language modeling and fine-tuning for 4 downstream tasks spanning document- and user-levels. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have led to the emergence of large language models capable ofvarious natural language processing tasks. |
| Approach: | They propose a multi-instructional training approach that integrates a large language model with a speech encoder to harness the capabilities of LLMs for speech recognition and beyond. |
| Outcome: | The proposed model can be trained and aligned with a multilingual LLM on 1900 hours of transcribed data from 139 languages. |
Copied to clipboard
| Challenge: | Existing Transformers models are computationally expensive for long context inputs. |
| Approach: | They propose a transformer that can interchange information between memory states and context . they evaluate the efficiency of their model on three dialogue datasets and two language datasets . |
| Outcome: | The proposed model is compatible with existing transformer models and can preserve dialogue history information. |
Copied to clipboard
| Challenge: | Recent advances in learning representations of visual and language information have been a problem with many applications. |
| Approach: | They propose to extract visual expressions from images aligned with linguistic expressions that describe the images to learn representations from implicit expressions. |
| Outcome: | The proposed representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition. |
Copied to clipboard
| Challenge: | Recent studies have proposed unified user modeling frameworks that leverage user behavior data from various applications. |
| Approach: | They propose to use user behavior sequences as plain text to represent rich information in any domain or system without losing generality. |
| Outcome: | The proposed frameworks achieve excellent results on diverse recommendation tasks and can be used on unseen domains and services. |
Copied to clipboard
| Challenge: | Existing methods to reduce run-times for language models with large word vocabularies are based on noise contrastive estimation (NCE) |
| Approach: | They propose to use noise-constrained noise-based models to approximate the normalized probability of a class without having to compute the partition function. |
| Outcome: | The proposed model outperforms softmax-based models in a variety of NLP tasks and is based on the noise-constrained noise-constant estimation properties. |
Copied to clipboard
| Challenge: | Existing pre-trained language models with self-attention encoder architectures are less useful in practice. |
| Approach: | They propose to use user and system tokens to model dialogue behavior during pre-training . they propose a contrastive objective function to simulate the response selection task . |
| Outcome: | The proposed model outperforms baseline models on four downstream tasks . it also has a few-shot ability that can mitigate the data scarcity problem . |
Copied to clipboard
| Challenge: | Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the researched class. |
| Approach: | They propose a set of NLP benchmarks for the Turkish language that contains several NLP tasks. |
| Outcome: | The proposed benchmarks outperform previous work significantly in the Turkish language. |
Copied to clipboard
| Challenge: | Recent work shows that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining. |
| Approach: | They propose a resource-efficient and modular domain specialization by means of domain adapters in which domain knowledge is encoded. |
| Outcome: | The proposed framework extracts domain-specific terms and then uses them to build DomainCC and DomainReddit resources based on masked language modeling and response selection objectives. |
Copied to clipboard
| Challenge: | Recurrent neural networks have achieved state-of-the-art results in many artificial intelligence tasks, such as language modeling, neural machine translation and speech recognition. |
| Approach: | They propose an efficient architecture to improve the efficiency of such RNN model training by adopting the group strategy for recurrent layers while exploiting the representation rearrangement strategy between layers as well as time steps. |
| Outcome: | The proposed architecture achieves comparable or better accuracy compared with baselines, with a much smaller number of parameters and at a lower computational cost. |
Copied to clipboard
| Challenge: | Recent advances in image tokenizers have enabled text-to-image generation using auto-regressive methods, but these methods lack pre-trained language models for text-based models. |
| Approach: | They adapt a pre-trained language model for auto-regressive text-to-image generation and show that pre-train language models offer limited help. |
| Outcome: | The proposed model is compared with a pre-trained language model and shows that it is no more effective than random initialized models. |
Copied to clipboard
| Challenge: | Existing RALM methods focus on modifying the LM architecture to facilitate incorporation of external information, complicating deployment. |
| Approach: | They propose to condition a language model on relevant documents from a grounding corpus during generation by conditioning on external knowledge sources. |
| Outcome: | The proposed method significantly improves language modeling performance and provides natural source attribution mechanism. |
Copied to clipboard
| Challenge: | Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measurements of harms. |
| Approach: | They apply a measurement modeling lens to inventory pitfalls that threaten benchmarks' validity as measurement models for stereotyping. |
| Outcome: | The proposed benchmarks lack clarity and assumptions that affect how they conceptualize and operationalize stereotyping. |
Copied to clipboard
| Challenge: | Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars. |
| Approach: | They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources. |
| Outcome: | The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios. |
Copied to clipboard
| Challenge: | Pre-trained language models have increased the performance of data-driven natural language processing (NLP) models on a wide variety of tasks. |
| Approach: | They propose a model-free approach to probing via prompting which formulates probing as a prompting task and combine pruning to analyze where the model stores the linguistic information in its architecture. |
| Outcome: | The proposed approach extracts information from pre-trained models while learning much less on its own. |
Copied to clipboard
| Challenge: | Recent models have added structure to recurrent neural networks at the cost of giving up exact inference, or using soft structure instead of latent variables. |
| Approach: | They propose a syntactic generative model with exact marginalization that supports dependency parsing and language modeling. |
| Outcome: | The proposed models achieve state-of-the-art for supervised dependency parsing and language modeling. |
Copied to clipboard
| Challenge: | Abstract meaning representation (AMR) is a semantic graph representation that abstracts meaning away from a sentence. |
| Approach: | They propose a decoder that back predicts projected AMR graphs on target sentences . their results show superiority over previous state-of-the-art decoded graph Transformer . |
| Outcome: | The proposed model outperforms the state-of-the-art model on two AMR benchmarks. |
Copied to clipboard
| Challenge: | Language models are a key component of natural language processing, but their size is a problem because they are typically trained with a closed output vocabulary derived from the training data. |
| Approach: | They propose a fully compositional output embedding layer for language models that is grounded in semantically related words and free-text definitions. |
| Outcome: | The proposed model outperforms state-of-the-art methods and adaptation approaches on cross-domain modeling and cross-learning tasks. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have resulted in powerful generation models, but their style is implicitly dependent on the training data and cannot emulate a specific target style. |
| Approach: | They propose an approach to induce certain target-author attributes by incorporating continuous multi-dimensional lexical preferences of an author into generative language models. |
| Outcome: | The proposed model generates text that aligns with a given target author’s lexical style and is competitive with baselines. |
Copied to clipboard
| Challenge: | Existing work on hierarchical structure in neural networks has not captured human intuitions about hierarchic structures. |
| Approach: | They propose to add an extra constraint to attention heads of the bidirectional Transformer encoder to encourage attention heads to follow tree structures. |
| Outcome: | The proposed model improves language modeling and learning more explainable attention scores. |
Copied to clipboard
| Challenge: | a new massive multilingual dataset is available for language modeling and machine translation training. |
| Approach: | They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora . |
| Outcome: | The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 . |
Copied to clipboard
| Challenge: | Federated Learning (FL) is a machine learning technique that trains a model across multiple distributed clients holding local data samples, without ever storing client data in a central location. |
| Approach: | They propose to use pretrained models to study three multilingual language tasks . they also examine impact of non-IID text on FL in naturally occurring data . |
| Outcome: | The proposed methods perform better than centralized learning even when using non-IID partitioning. |
Copied to clipboard
| Challenge: | Existing methods to combine language modeling and knowledge graphs (KG) lack the context to provide a more precise understanding of the concepts. |
| Approach: | They propose to use external entity descriptions to provide contextual information for commonsense question answering models. |
| Outcome: | The proposed model achieves state-of-the-art among non-generative models in OpenBookQA and is the first of its kind in the field. |
Copied to clipboard
| Challenge: | a new approach to natural language processing uses arbitrary symbols to represent meaning . Soundex, MetaPhone, NYSIIS, logogram are used as inputs for NLP . |
| Approach: | They propose to use arbitrary symbols to represent linguistic meaning of a word . they propose to integrate codewords with text to provide more reliable inputs . |
| Outcome: | The proposed approach outperforms state-of-the-art models on machine translation, language modeling, and part-of speech tagging. |
Copied to clipboard
| Challenge: | Recent advances in the field of language modeling have improved state-of-the-art results on many natural language processing tasks. |
| Approach: | They propose to use a French Question Answering Dataset to track progress of French Question answering models. |
| Outcome: | The proposed model achieves an F1 score of 92.2 and an exact match ratio of 82.1 on the test set. |
Copied to clipboard
| Challenge: | Recent work on retrieval-augmented language models has shown impressive results . performance gains from retrieval to a large extent originate from overlapping tokens between the database and test data, suggesting less of non-trivial generalization than previously assumed. |
| Approach: | They propose to off-load memory from trainable weights to a retrieval database and compare it to larger models with a larger model. |
| Outcome: | The proposed model outperforms GPT-3 and Jurassic-1 on the Pile at 4% of the model parameters. |
Copied to clipboard
| Challenge: | Debiasing Pretrained Language Models (PLMs) are task-agnostic and can be generalizable, but its impact on language modeling ability and the risk of relearning social biases remain as the two most significant challenges. |
| Approach: | They propose a framework which can Propagate Socially-fair Debiasing to Downstream Fine-tuning to alleviate the forgetting issue of PLMs by regularizing debiased attention heads based on the PLM’s bias levels from stages of pretraining and debiase. |
| Outcome: | The proposed framework can Propagate Socially-fair Debiasing to Downstream Fine-tuning, indicating that the ineffectiveness of debiase can be alleviated by overcoming the forgetting issue through regularizing successfully debiased attention heads based on the PLMs’ bias levels from stages of pretraining and debiases. |
Copied to clipboard
| Challenge: | Experimental results show that n-gram models can achieve satisfactory performance on a large proportion of testing cases. |
| Approach: | They propose to learn a neural LM that fits the residual between an n-gram LM and the real-data distribution. |
| Outcome: | The proposed model achieves additional performance gains over popular standalone models on three typical language tasks. |
Copied to clipboard
| Challenge: | Initial dropout was seen as a breakthrough regularization technique that reduced overfitting, yet single-epoch pretraining tasks common to modern LLMs yield minimal overfit. |
| Approach: | They propose to use dropout during single-epoch pretraining to reduce overfitting in language modeling, morpho-syntax, question answering, and MNLI to improve performance. |
| Outcome: | The results show that dropout is not used in large LLMs and improves performance in language modeling, morpho-syntax, question answering, and MNLI. |
Copied to clipboard
| Challenge: | RNNGs model syntax and structure by incrementally generating a syntax tree and sentence in a top-down, left-to-right order. |
| Approach: | They explore unsupervised learning of recurrent neural network grammars for language modeling and grammar induction. |
| Outcome: | The proposed model outperforms standard sequential language models and improves parsing performance. |
Copied to clipboard
| Challenge: | Recent studies show that self-attention based models have limitations on modeling sequential transformations. |
| Approach: | They propose to extract some explainable features from trained RNNs that are reminiscent of classical n-grams features. |
| Outcome: | The proposed models can model interesting linguistic phenomena such as negation and intensification. |
Copied to clipboard
| Challenge: | Transformers are impressive but inefficient and costly, which limits their applications and accessibility. |
| Approach: | They first use different ways to downsample and upsamplify activations in Transformers to make them hierarchical. |
| Outcome: | The proposed model outperforms Transformers on the ImageNet32 and enwik8 benchmarks. |
Copied to clipboard
| Challenge: | Current language models are unable to efficiently model entity names observed in text providing insufficient context. |
| Approach: | They propose to augment a traditional model with an external knowledge base to model entity names observed in text. |
| Outcome: | The proposed model improves on a Named Entity Recognition (NER) task by requiring no additional information such as named entity tags. |
Copied to clipboard
| Challenge: | Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities. |
| Approach: | They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness . |
| Outcome: | Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds . |
Copied to clipboard
| Challenge: | Existing methods to evaluate reliability of generated text are lacking in natural language generation. |
| Approach: | They propose a non-exchangeable conformal prediction method that provides bounds on coverage . they validated their method with k-NN retrieval and show that it produces encouraging results . |
| Outcome: | The proposed method produces encouraging results in machine translation and language modeling tasks. |
Copied to clipboard
| Challenge: | Failing to capture the structure of input language could lead to generalization problems and over-parametrization. |
| Approach: | They propose a new syntax-aware language model that explicitly models the structure with an incremental parser and maintains the conditional probability setting of a standard language model. |
| Outcome: | The proposed model can achieve strong results in language modeling, parsing, and syntactic generalization tests while using fewer parameters than other models. |
Copied to clipboard
| Challenge: | Recent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on. |
| Approach: | They propose to use Counterfactual Data Augmentation, Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebia as bias mitigation techniques to quantify their effectiveness. |
| Outcome: | The proposed techniques are Counterfactual Data Augmentation (CDA), Dropout, Iterative Nullspace Projection, Self-Debias, and SentenceDebia. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have made it viable to model language as distributions over characters. |
| Approach: | They propose to leverage internal states of a trained character language model to produce a new type of word embeddings. |
| Outcome: | The proposed embeddings outperform the state-of-the-art on four classic sequence labeling tasks. |
Copied to clipboard
| Challenge: | Recent work shows that training text encoders using data from multiple tasks helps to produce an encoder that can be used in numerous downstream tasks with minimal fine-tuning. |
| Approach: | They incorporate four different tasks to improve abstractive summarization performance . they use a pretrained BERT model and train all tasks using a small-scale training corpus . |
| Outcome: | The proposed model outperforms a model trained in a multitask setting with no additional summarization data. |
Copied to clipboard
| Challenge: | Existing methods to improve language modeling performance are based on regularized LSTMs with a large number of parameters and training time. |
| Approach: | They propose a method that decodes the last token in context using the predicted distribution of the next token. |
| Outcome: | The proposed method improves perplexity on the Penn Treebank dataset by 1.8 points and 2.3 points on the WikiText-2 datasets. |
Copied to clipboard
| Challenge: | Existing long-context models degenerate with retrieved contexts. |
| Approach: | They propose a framework that can be applied to existing decoder-only LLMs for context expansion. |
| Outcome: | The proposed framework can be applied to any existing decoder-only LLMs for context expansion. |
Copied to clipboard
| Challenge: | Common language models typically predict the next word given a past context. |
| Approach: | They propose a method that aligns the given context and the following phrase . they define syntactic heights and phrase segmentation rules to enable it to learn . |
| Outcome: | The proposed model outperforms strong baseline models on Wikitext-103 dataset. |
Copied to clipboard
| Challenge: | masked language models are trained on ever larger corpora, but pre-training on a modestly-sized but representative, well-balanced, and publicly available corpus can reach better performance than the original BERT model. |
| Approach: | They propose an optimized LM architecture called LTG-BERT that can be used to train a competitive language model on a small and standardizable corpus. |
| Outcome: | The proposed architecture outperforms the original English BERT model on a representative, well-balanced and publicly available corpus. |
Copied to clipboard
| Challenge: | Recent studies show that neural models lack strong intuitions . recent studies show connections between convolutional neural networks and weighted finite state automata (WFSAs) |
| Approach: | They show that some recurrent neural networks share a connection to weighted finite state automata (WFSAs) they define rational recurrences as recursive hidden state update functions . they propose to use these functions to write forward calculations of a finite set of WFSA's . |
| Outcome: | The proposed model outperforms two baselines on language modeling and text classification. |
Copied to clipboard
| Challenge: | Pretrained language models often need to specialize to specific domains. |
| Approach: | They propose an approach that performs weight-space averaging of adapters trained on different domains. |
| Outcome: | The proposed approach improves performance to new domains without extra training. |
Copied to clipboard
| Challenge: | Inflected words benefit more from explicitly modeling morphology than uninflectes . morphological supervision is also used to augment character language models in low-resource languages . |
| Approach: | They add morphological supervision to character language models via multitasking to improve BPC performance across 24 languages even when morphology data and language modeling data are disjointed. |
| Outcome: | The addition improves performance even when morphology data and language modeling data are disjointed. |
Copied to clipboard
| Challenge: | Language models rely on massive web crawls for diverse text data, but are rife with undesirable content. |
| Approach: | They analyze newspaper articles written by students from across the country to determine whose language is preferred by a quality filter. |
| Outcome: | The results show that newspapers from wealthier, educated, and urban zones are more likely to be classified as high quality. |
Copied to clipboard
| Challenge: | Question answering (QA) tasks have been posed using a variety of formats . a new study aims to develop specialized QA models that can be used to train QA systems . |
| Approach: | They build a pre-trained question answering model that performs well across 19 QA datasets . they argue that format-specialized models can limit the ability to teach reasoning . |
| Outcome: | a new model that trains on QA datasets performs on par with 8 models trained on individual datasets . a single model that trained on UNIFIEDQA performs well on 19 QA data . |
Copied to clipboard
| Challenge: | Existing conditional text generation models produce unfaithful and unfaithed summaries . current models accomplish a high level of fluency and coherence . |
| Approach: | They propose to use pretrained models for document summarization to better understand hallucinations . they find that textual entailment measures better correlate with faithfulness . |
| Outcome: | The proposed models generate faithful and factual summaries as evaluated by humans. |
Copied to clipboard
| Challenge: | pixel-based language modeling integrates visual and textual data to improve performance of language models. |
| Approach: | They propose a method that integrates visual and textual data into an autoregressive framework. |
| Outcome: | The proposed method improves performance of pixel-based language models by incorporating visual and textual data. |
Copied to clipboard
| Challenge: | Simile is a special type of metaphor, where comparators such as like and as are used to compare two objects. |
| Approach: | They propose a neural network framework for simile sentence classification, simile component extraction and language modeling. |
| Outcome: | The proposed framework outperforms rule-based and feature-based approaches in simile sentence classification and simile component extraction tasks. |
Copied to clipboard
| Challenge: | Recent studies have shown that architecture search can improve performance on language modeling and image classification tasks with reasonable training speed. |
| Approach: | They propose a continual architecture search approach that continually evolves the model parameters during sequential training of several tasks without losing performance on previously learned tasks. |
| Outcome: | The proposed approach improves language modeling and image classification with reasonable training speed and a weight-sharing strategy. |
Copied to clipboard
| Challenge: | Natural language processing is one of the most important fields of artificial intelligence. |
| Approach: | They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens . |
| Outcome: | The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles . |
Copied to clipboard
| Challenge: | Existing long context models suffer from performance decline when the input text exceeds their length limit. |
| Approach: | They propose a multi-task long context benchmark to evaluate LLMs' long context ability using 10 datasets from 5 different NLP tasks. |
| Outcome: | The proposed model covers 5 domains and core capacities of large language models. |
Copied to clipboard
| Challenge: | a lack of multilingual multimodal datasets has hindered multimodal vision and language modeling efforts. |
| Approach: | They propose a multilingual evaluation benchmark for the visual question answering task . they extend the established English GQA dataset to 7 typologically diverse languages . |
| Outcome: | The proposed methods outperform current state-of-the-art models in zero-shot cross-lingual settings, but the accuracy remains low across languages. |
Copied to clipboard
| Challenge: | Speculative sampling is an efficient way to accelerate the auto-regressive generation process of large language models. |
| Approach: | They propose a frequency-ranked speculative sampling framework that optimizes draft candidate selection through vocabulary space compression. |
| Outcome: | Experiments show that FR-Spec reduces LM Head computation overhead by 75% while ensuring the equivalence of the final output distribution. |
Copied to clipboard
| Challenge: | Existing methods to protect sensitive data from leaking are over-pessimistic and undifferentiated. |
| Approach: | They propose a new privacy notion, selective differential privacy, to provide rigorous privacy guarantees on the sensitive portion of the data to improve model utility. |
| Outcome: | The proposed privacy-preserving mechanism achieves better utility while remaining safe under various privacy attacks compared to baselines. |
Copied to clipboard
| Challenge: | a growing demand for the ability to communicate in English means automated tutoring and assessment systems are becoming more popular. |
| Approach: | They propose to use automatic speech recognition transcripts to grade spontaneous speech based on textual features. |
| Outcome: | The proposed system improves on a transformer encoder with native language identification as an auxiliary task. |
Copied to clipboard
| Challenge: | Subword tokenization algorithms have been an essential component of language modeling but their static nature results in important flaws that degrade the models’ downstream performance and robustness. |
| Approach: | They propose a module for Adaptive Neural TokenizAtion that is differentiable and trained end-to-end with the language model. |
| Outcome: | The proposed tokenizer improves robustness to character perturbations and out-of-domain data. |
Copied to clipboard
| Challenge: | Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens. |
| Approach: | They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions. |
| Outcome: | The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000. |
Copied to clipboard
| Challenge: | Modern machine learning works with massive amounts of data on a range of tasks like language modeling, object detection, and data mining. |
| Approach: | They propose a probabilistic robustness rewarded data optimization approach to enhance the model's generalization power by selecting training data that optimizes probabilistic metrics. |
| Outcome: | The proposed approach achieves +17.2% increase of accuracy and -28.05 decrease of perplexity on unknown-domain test sets. |
Copied to clipboard
| Challenge: | Entity Matching (EM) aims at recognizing entity records that denote the same real-world object. |
| Approach: | They propose a novel EM framework that consists of Heterogeneous Information Fusion and Key Attribute Tree Induction to decouple feature representation from matching decision. |
| Outcome: | The proposed framework outperforms SOTA EM models on 6 public datasets and 3 industrial datasets. |
Copied to clipboard
| Challenge: | Existing methods to verify scientifically false online information are limited by the lack of training data in the scientific domain. |
| Approach: | They propose an in-domain language modeling method for fact extraction and verification systems . they use SCIFACT to extract scientifically false online information . |
| Outcome: | The proposed method improves accuracy 30% on SCIFACT dataset . state-of-the-art model achieves only 46.6% precision, which is hard to be trusted for users. |
Copied to clipboard
| Challenge: | Existing word-based approaches to learning word representations are blind to subword information in words. |
| Approach: | They propose a character-based word representation approach to learn word representations from characters. |
| Outcome: | The proposed model outperforms baseline models that regard words as atomic units . the proposed model achieves 18.5% improvement on average in perplexity for morphologically rich languages . |
Copied to clipboard
| Challenge: | Existing methods for LRM unlearning overlook critical information leakage in reasoning traces, even when final answers are successfully removed. |
| Approach: | They propose a method that suppresses reasoning traces while preserving the model's general reasoning ability. |
| Outcome: | The proposed method significantly reduces reasoning trace leakage and achieves strong performance across reasoning and safety benchmarks, including WMDP, StrongReject, JBB-Behaviors and WildJailbreak. |
Copied to clipboard
| Challenge: | Using adversarial triggers, a model can produce a specific prediction . adversarial attacks are useful for evaluation and interpretation . |
| Approach: | They propose a gradient-guided search over tokens that finds short adversarial triggers that successfully trigger the target prediction. |
| Outcome: | The proposed algorithm finds short trigger sequences that successfully trigger the target prediction. |
Copied to clipboard
| Challenge: | Existing research on cultural background modeling is coarse-grained and does not examine cultural differences among speakers of the same language. |
| Approach: | They use a news-based cultural background prediction dataset to annotate, validate and benchmark NLP models with cultural background features. |
| Outcome: | The proposed model improves on nine syntactic, semantic, and psycholinguistic tasks while introducing cultural background information does not improve the Go-Emotions task due to text domain conflicts. |
Copied to clipboard
| Challenge: | Infilling is the task of predicting missing spans of text at any position in a document. |
| Approach: | They propose a framework which can be used to infill entire sentences . they train off-the-shelf LMs on sequences containing concatenation of masked text . |
| Outcome: | The proposed approach can infill entire sentences on short stories, scientific abstracts, and lyrics. |
Copied to clipboard
| Challenge: | Experimental results show that REtrieving from the traINing datA only can lead to significant gains on multiple NLG and NLU tasks. |
| Approach: | They propose to retrieve training instances from traINing datA and concatenate them with input to generate output. |
| Outcome: | The proposed method achieves state-of-the-art results on XSum, BigPatent, and CommonsenseQA. |
Copied to clipboard
| Challenge: | Existing models for document-level language pretraining are not suitable for long documents due to their quadratically increasing memory and time consumption. |
| Approach: | They propose a document-level language pretraining model based on Recurrence Transformers. |
| Outcome: | The proposed model outperforms existing models on language understanding tasks. |
Copied to clipboard
| Challenge: | Pretrained Transformer encoders are the dominant approach to sequence labeling . however, few have been applied to sequence labels on flat or simplified tasks . |
| Approach: | They propose to use pretrained Transformer encoders to model relations across words . they find that the architectures adapt well across tagging tasks that vary in complexity . |
| Outcome: | The proposed architectures perform well across tagging tasks across languages and datasets. |
Copied to clipboard
| Challenge: | Transformer-based language models have a finite context window and expensive computational cost of processing long text documents. |
| Approach: | They propose to adapt pre-trained LMs into AutoCompressors to compress text into summary vectors . authors propose to use summary vector to speed up inference over long contexts based on a finite context window . |
| Outcome: | The proposed model can compress long contexts into summary vectors, which are accessible as soft prompts. |
Copied to clipboard
| Challenge: | Existing studies have explored compression and accumulation methods to compress contexts, but these methods lose useful context information during the compression process, leading to performance degradation. |
| Approach: | They propose a method that allows LLMs to take a deep breath and insert a special token at the end of each chunk. |
| Outcome: | Experiments on language modeling and out-of-domain tasks validate the superiority of the proposed method. |
Copied to clipboard
| Challenge: | Variational Autoencoder (VAE) is widely used to approximate a model’s posterior on latent variables. |
| Approach: | They propose to let the Kullback–Leibler divergence individual follow a distribution across the whole dataset and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL’s distribution positive. |
| Outcome: | The proposed approach can avoid posterior collapse effectively and efficiently without introducing any new model component or modifying the objective. |
Copied to clipboard
| Challenge: | Multilingual language models are widely used to extend NLP systems to low-resource languages. |
| Approach: | They pre-train over 10,000 monolingual and multilingual language models for over 250 languages including multiple language families that are under-studied in NLP. |
| Outcome: | The results show that adding multilingual data improves low-resource language modeling performance, similar to increasing low-source dataset sizes by up to 33%. |
Copied to clipboard
| Challenge: | Existing studies on the definitions of good privacy for natural language use argue that different applications and models require different definitions. |
| Approach: | They propose a technique that uses a combination of statistics and language modeling to produce high (768) dimensional, general -SentDP document embeddings that guarantee a single sentence can be substituted with any other sentence. |
| Outcome: | The proposed method outperforms baseline methods with weaker guarantees like word-level Metric DP and outperformed baseline methods. |
Copied to clipboard
| Challenge: | Word-embeddings are vital components of natural language processing (NLP) but they consume a lot of memory which poses a challenge for edge deployment. |
| Approach: | They propose an embedding compression method based on matrix decomposition and knowledge distillation that initializes weights of pre-trained word-embeddings and fine-tunes end-to-end. |
| Outcome: | The proposed method has higher BLEU score on translation and lower perplexity on language modeling compared to complex, difficult to tune methods. |
Copied to clipboard
| Challenge: | Recent approaches to architecture search have shown good improvements in terms of performance with reasonable training speed. |
| Approach: | They propose an algorithm with more activation functions, input edges, and atomic operations to search for architectures that are optimal for given task. |
| Outcome: | The proposed algorithm reproduces well-known LSTM and GRU architectures and initializes with them for finding architectures more efficiently. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are rapidly deployed and continue to evolve through scaling. |
| Approach: | They propose a method to train strong long-context LLMs that are capable of utilizing massive context windows of up to 32,000 tokens. |
| Outcome: | The proposed model can surpass gpt-3.5-turbo-16k's overall performance on long-context benchmarks with a cost-effective instruction tuning procedure that is free of expensive annotations. |
Copied to clipboard
| Challenge: | ConVEx is an efficient pretraining and fine-tuning neural approach for slot-labeling dialog tasks. |
| Approach: | They propose an efficient pretraining and fine-tuning neural approach for slot-labeling dialog tasks that uses a pairwise cloze task and reddit data. |
| Outcome: | The proposed approach is well aligned with its intended use on slot-labeling tasks and can be used across a range of domains and data sets. |
Copied to clipboard
| Challenge: | Traditional NLP has long held (supervised) syntactic parsing necessary for successful higher-level semantic language understanding (LU). |
| Approach: | They empirically examine the usefulness of supervised parsing for semantic LU in LM-pretrained transformer networks. |
| Outcome: | The proposed model is based on LM-pretrained transformer networks with a biaffine parsing head and fine-tuned for LU tasks. |
Copied to clipboard
| Challenge: | Existing tokenization methods focus on information-theoretical goals like high compression and low fertility rather than linguistic goals like morphological alignment. |
| Approach: | They propose to incorporate morphological knowledge into tokenization to improve both morphology and downstream performance. |
| Outcome: | The proposed tokenization improves overall performance on four downstream tasks. |
Copied to clipboard
| Challenge: | Hierarchical Multiscale LSTM model learns structure from character input . high complexity of architecture, training and implementations might hinder its applicability . |
| Approach: | They propose to reproduce and ablate hierarchical multiscale LSTM language model and show that simplifying certain aspects of the architecture can improve its performance. |
| Outcome: | The proposed model performs better when simplified and linguistic units are learned by different levels of the model. |
Copied to clipboard
| Challenge: | Existing approaches to align large language models with human values and preferences are not able to be applied to all tasks and fields. |
| Approach: | They propose a high-dimensional representation of symbolic human value distributions in LLMs that is orthogonal to model architecture and training data. |
| Outcome: | The proposed representations are evaluated on 15 open-source and commercial LLMs and are self-supervised from the value-relevant output of 8 LLM models. |
Copied to clipboard
| Challenge: | a current paradigm of language modeling discards linguistic relations between tokens during tokenization, creating a fundamental gap . empirical results show that TriEmbed provides more linguistically informative token embeddings . |
| Approach: | They propose a reparameterization method that incorporates morphological relationships . they propose to organize the vocabulary into a Trie structure to reparametrize embeddings . |
| Outcome: | Empirical results show that TriEmbed outperforms existing token embeddings while offering more linguistically informative token embeds. |
Copied to clipboard
| Challenge: | a study examines whether readers can distinguish between two types of reading goals: information seeking and ordinary reading for comprehension. |
| Approach: | They propose a method to distinguish between two types of reading goals: information seeking and ordinary reading for comprehension. |
| Outcome: | The proposed model solves the reading goal-oriented task with the most accurate predictions in real time, the authors say . |
Copied to clipboard
| Challenge: | Semi-parametric models augment generation with retrieval, but require expensive retrieval operation for every generated token. |
| Approach: | They propose a semi-parametric model which augments generation with retrieval by retrieving tokens from a datastore. |
| Outcome: | The proposed model can retrieve chunks of tokens from the datastore, instead of a single token, with a low decoding speed. |
Copied to clipboard
| Challenge: | Term memory networks (RNNs) are difficult to optimize due to gradient vanishing and explosion. |
| Approach: | They propose a neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. |
| Outcome: | The proposed method improves state-of-the-art performance on short and long sequences and generates coherent, novel text articles with thousands of tokens. |
Copied to clipboard
| Challenge: | Existing diffusion models have limitations in modeling discrete data, e.g., languages . we present a novel diffusion model for language modeling inspired by linguistic features in languages based on iterative denoising . |
| Approach: | They propose a method that iteratively denoises and adds corruptions to the textual data through soft-masking to better noise it. |
| Outcome: | The proposed model achieves better generation quality and lower training cost than current models with better performance. |
Copied to clipboard
| Challenge: | Using word-based models, we compare word-oriented models with char-based ones . word-driven models are more vulnerable to data sparsity and the presence of out-of-vocabulary words . |
| Approach: | They benchmark word-based models with char-based model which does not involve word segmentation in four NLP benchmark tasks. |
| Outcome: | The proposed model outperforms char-based models in four NLP benchmark tasks. |
Copied to clipboard
| Challenge: | Existing Diffusion Language Models rely on hard binary masking and discrete token assignments, which hinder the revision of early decisions. |
| Approach: | They propose a diffusion-based language modeling approach that replaces hard binary masks with evolving soft token distributions. |
| Outcome: | The proposed approach outperforms existing DLMs on multiple benchmarks. |
Copied to clipboard
| Challenge: | Distillation efforts have led to language models that are more compact and efficient without serious drops in performance. |
| Approach: | They propose to augment distillation with a third objective that encourages the student model to imitate the causal dynamics of the teacher through a distillation interchange intervention training objective (DIITO). |
| Outcome: | The proposed method lowers perplexity on the WikiText-103M corpus and improves on the GLUE benchmark, SQuAD, and CoNLL-2003. |
Copied to clipboard
| Challenge: | Existing models that use transformers to model language cost quadratically increase with sequence length. |
| Approach: | They propose a selective cache which stores key-value pairs from previous contexts. |
| Outcome: | The proposed selective cache outperforms XL cache and compressive cache by considerable margins. |
Copied to clipboard
| Challenge: | Recent Open Information Extraction systems allow us to extract ever larger (yet incomplete) open-domain Knowledge Bases from text. |
| Approach: | They propose a baseline model which gives competitive results in a previously defined protocol and provides an independent source of signal to judge arbitrary fact plausibility. |
| Outcome: | The proposed model gives competitive results in the previously defined protocol and provides an independent source of signal to judge arbitrary fact plausibility. |
Copied to clipboard
| Challenge: | SANs are an integral part of successful neural networks such as Transformer . training SAN on a task or pretraining them on language modeling requires large amounts of data and compute resources. |
| Approach: | They propose to modify SANs to enable faster learning, i.e., higher accuracies after fewer update steps. |
| Outcome: | The proposed modifications enable faster learning, i.e., higher accuracies after fewer update steps. |
Copied to clipboard
| Challenge: | Existing evaluation metrics poorly approximate parser quality, says a new study . questions under discussion is a linguistic framework that views discourse as asking questions and answering them . |
| Approach: | They propose a framework for automatic evaluation of QUD parsing . they use a dataset of fine-grained evaluation of 2,190 QUD questions . |
| Outcome: | The proposed framework shows that satisfying constraints of QUD is still challenging for modern LLMs. |
Copied to clipboard
| Challenge: | Existing models of language understanding are based on explicit representations of hierarchical structure, but there are good reasons to doubt that they can be said to understand language in any meaningful way. |
| Approach: | They examine whether syntactic and semantic graph representations can complement and improve neural language modeling. |
| Outcome: | The proposed model outperforms pretrained models on English WSJ in perplexity and other metrics. |
Copied to clipboard
| Challenge: | Speculative decoding is a widely used technique to speed up inference for Large Language Models (LLMs) Autoregressive decoding has been known to be hardware inefficient, leading to poor resource utilization and low throughput during inference. |
| Approach: | They propose to use a draft model to generate speculative tokens and then use the target LLM to verify those tokens. |
| Outcome: | The proposed model can provide 111% higher throughput than existing draft models and generalizes further to all LLaMA models and supervised fine-tuned models. |
Copied to clipboard
| Challenge: | Despite the small pool of speakers, there are still few natural language processing corpora, studies or tools for Swiss German. |
| Approach: | They propose to use a web scraper to generate the largest Swiss German text corpus . they show that the tool can be applied to other low-resource languages as well . |
| Outcome: | The proposed tool significantly improves language modeling in Swiss German, the authors show . |
Copied to clipboard
| Challenge: | Word embeddings are usually derived from corpora containing text from many individuals . however, they cannot account for user-specific word preferences, such as using the same word in different ways across contexts. |
| Approach: | They propose a new form of personalized word embeddings that use demographic-specific word representations derived compositionally from full or partial demographic information for a user. |
| Outcome: | The proposed representations outperform generic representations on two English language tasks. |
Copied to clipboard
| Challenge: | Existing work on answer-aware questions generates a sentence and answer span as input . previous work on QG was mainly tackled by rule-based approach and neural-based one . |
| Approach: | They propose to incorporate an auxiliary task of language modeling to help question generation in a hierarchical multi-task learning structure. |
| Outcome: | The proposed model improves on SQuAD and MARCO datasets and human evaluation proves it. |
Copied to clipboard
| Challenge: | Unsupervised parsing is a form of reinforcement learning that improves syntactic structures but lacks interpretability due to its lack of ad hoc heuristics. |
| Approach: | They propose an unsupervised approach that transfers syntactic knowledge to a Tree-LSTM model with discrete parsing actions. |
| Outcome: | The proposed model outperforms existing models on the All Natural Language Inference dataset and achieves a new state of the art in terms of parsing F-score. |
Copied to clipboard
| Challenge: | Recent improvements in NLP tasks can be attributed to the Transformer model. |
| Approach: | They propose to use parameter-sharing methods to reduce parameter budgets in generative models by using sandwich-style parameter sharing and self-attentive embedding factorization. |
| Outcome: | The proposed model outperforms the current RNN model even with significantly fewer parameters. |
Copied to clipboard
| Challenge: | Existing question answering datasets provide extractive or short answers, but less attention has been paid to open-ended questions that require explanations. |
| Approach: | They present a large-scale corpus for long form question answering . they use a Reddit forum to provide elaborate answers to open-ended questions . |
| Outcome: | The proposed model outperforms Seq2Seq, language modeling, and other models in human evaluations. |
Copied to clipboard
| Challenge: | Recent sparse decoding methods improve efficiency but suffer from KV cache misalignment, resulting in performance degradation. |
| Approach: | They propose a method that combines block-sparse attention with periodic dense rectification to bound error accumulation and preserve alignment with the pretraining distribution. |
| Outcome: | Experiments on math reasoning, language modeling, and retrieval tasks show that ReSA achieves near-lossless generation quality with significantly improved efficiency. |
Copied to clipboard
| Challenge: | Syntactic language models (SLMs) incorporate syntactical biases into Transformers . authors identify key aspects of design choices in existing models and novel variants based on experimental results . |
| Approach: | They propose a framework that incorporates existing and new SLMs to enhance Transformers by incorporating syntactic biases. |
| Outcome: | The proposed framework improves on existing models and novel variants across language modeling, syntactic generalization, summarization, and inference efficiency. |
Copied to clipboard
| Challenge: | Natural language consists of implicit and underspecified phrases, which can cause misunderstandings. |
| Approach: | They propose to use wikiHow to extract human clarifications that resolve an implicit or underspecified phrase. |
| Outcome: | The proposed model can be used to generate alternate clarifications, which may or may not be compatible with the human clarification. |
Copied to clipboard
| Challenge: | This study examines the ability of Large Language Models to encapsulate cultural nuances across diverse linguistic landscapes. |
| Approach: | They examine the efficacy of language-specific instruction tuning and the impact of pretraining on dominant language data in Large Language Models. |
| Outcome: | The findings highlight a nuanced landscape, with inconsistencies and biases, particularly in non-Western cultures. |
Copied to clipboard
| Challenge: | Existing studies show that multilingual transformers are less effective in resource-lean scenarios and for distant languages. |
| Approach: | They propose to use massively multilingual transformers to pretrain languages . they show that MMTs are less effective in resource-lean scenarios and distant languages if they are pre-trained via language modeling . |
| Outcome: | The proposed model is less effective in resource-lean scenarios and for distant languages than cross-lingual word embeddings. |
Copied to clipboard
| Challenge: | Agentic learning increasingly hinges on interaction, yet real-world experience is expensive, limited, and often irreversible at inference time. |
| Approach: | They propose a framework that reframes language modeling as next-state prediction under interaction. |
| Outcome: | The proposed framework evaluates world models in text-based environments . it shows that sufficiently trained models capture coherent environment dynamics . |
Copied to clipboard
| Challenge: | Neural architecture search (NAS) is a popular approach for finding new models and freeing researchers from the hard work of designing network architectures. |
| Approach: | They propose differentiable neural architecture search methods for natural language processing . they remove the softmax-local constraint and apply it to named entity recognition . |
| Outcome: | The proposed method outperforms strong baselines on the language modeling task. |
Copied to clipboard
| Challenge: | morphological differences between languages are unclear, but are often considered unimportant . confounding factors make it hard to compare results and draw conclusions, authors argue . |
| Approach: | They propose to use token bigram metrics to predict difficulty of causal language modeling . they argue that confounding factors are contributing to the conflicting evidence . |
| Outcome: | The proposed metrics better capture the relation between morphology and tokenization compared to word-based models. |
Copied to clipboard
| Challenge: | Standard pre-trained language models do not see the characters that compose each token's string representation. |
| Approach: | They probe the embedding layer of pretrained language models and show that models learn the internal character composition of whole word and subword tokens without seeing the characters coupled with the tokens. |
| Outcome: | The embedding layers of RoBERTa and GPT2 hold enough information to accurately spell up to a third of the vocabulary and reach high character ngram overlap across all token types. |
Copied to clipboard
| Challenge: | Several efficient transformers have been proposed, but they all have a finite memory capacity and are forced to drop old information. |
| Approach: | They propose an unbounded long-term memory extension that extends the vanilla transformer by using a continuous-space attention mechanism to attend over the long-time memory. |
| Outcome: | The proposed model can model arbitrarily long contexts while keeping the computation budget fixed. |
Copied to clipboard
| Challenge: | Recent language models have shown strong data-fitting performance, but do not explicitly encode any notion of structural information. |
| Approach: | They propose a hybrid parser and neural language model that adds an attention layer over text spans in the left context. |
| Outcome: | The proposed model outperforms baseline models on language modeling and provides syntactically-informed representations of the context. |
Copied to clipboard
| Challenge: | Existing models with stacked layers do not explicitly model hierarchical structure of language understanding. |
| Approach: | They propose a recursive Transformer model based on differentiable CKY style binary trees to emulate hierarchical composition process. |
| Outcome: | The proposed model can predict words given their left and right abstraction nodes. |
Copied to clipboard
| Challenge: | Recent studies on large-scale in-context language models have reported successful in-const zero- and few-shot learning ability. |
| Approach: | They investigate the effects of the pretraining corpus on in-context learning in a Korean-centric model. |
| Outcome: | The study shows that pretraining corpus size does not determine in-context learning ability . the findings suggest that in-constext learning is not always competitive . |
Copied to clipboard
| Challenge: | LLaMa-based language model for materials science is first of its kind in the world . |
| Approach: | They propose an instruction-based process for trustworthy data curation in materials science (MatSci-Instruct) they then apply this process to finetune a LLaMa-based language model targeted for materials science. |
| Outcome: | The proposed model outperforms existing language models on materials science tasks and improves in successive stages of refinement. |
Copied to clipboard
| Challenge: | n-gram smoothing techniques were used to overcome overfitting problems in neural language models for decades. |
| Approach: | They propose to convert any n-gram smoothing technique into a regularizer compatible with neural language models. |
| Outcome: | The proposed regularizers outperform label smoothing on language modeling and machine translation. |
Copied to clipboard
| Challenge: | Attention pruning techniques have been developed to identify and exploit sparseness . previous work has taken pioneering steps to discover and explain the sparsity in attention patterns . |
| Approach: | They propose a framework that observes attention patterns in a fixed dataset and generates a global sparseness mask. |
| Outcome: | The proposed approach saves 90% of computations and maintains quality of results. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel on English NLU tasks, yet struggle to extend their NLU capabilities to underrepresented languages. |
| Approach: | They integrate machine translation models (MT) directly into LLM backbones via sample-efficient self-distillation. |
| Outcome: | The proposed model outperforms translation-test models on 127 low-resource languages. |
Copied to clipboard
| Challenge: | Recent work shows language models trained on form can capture aspects of meaning without explicit state supervision. |
| Approach: | They propose to use probing to "bake" state knowledge into language models . they propose to probe for underlying world state knowledge via text prompts . |
| Outcome: | The proposed methods show that language models trained on form can capture the world state without state supervision. |
Copied to clipboard
| Challenge: | Traditionally, researchers used manual coding to track conflict processes worldwide, but the high costs and slow pace of domain experts make it difficult and costly to monitor complex and rapidly changing conflicts. |
| Approach: | They propose a domain-specific pre-trained language model for conflict and political violence that can be used to train a language model from scratch and continue training. |
| Outcome: | The proposed model outperforms BERT in conflict research. |
Copied to clipboard
| Challenge: | atypical animacy is the property of being alive, but discrepancies are not uncommon . a typical animate is represented as either animate or inanimate in a text . |
| Approach: | They propose a method for determining whether an entity is represented as animate in a text . they use a nineteenth-century English text to analyze animacy . |
| Outcome: | The proposed method improves on an established animacy dataset and a newly introduced resource. |
Copied to clipboard
| Challenge: | In this paper, we test the hypothesis that deeper transformers generalize more compositionally. |
| Approach: | They propose to add layers to transformers to generalize more compositionally . they propose to fine-tune the models so that the total number of parameters is constant . |
| Outcome: | The proposed model generalizes more compositionally than shallower models, but returns diminish . the proposed model can be made shallower without sacrificing performance . |
Copied to clipboard
| Challenge: | Large language models have demonstrated their capability with few-shot inference . however, in-domain demonstrations are not always available in real scenarios . |
| Approach: | They propose unsupervised domain adaptation problem to adapt language models from source domain to target domain without any target labels. |
| Outcome: | The proposed model performs better than baseline models on Sentiment Analysis and Named Entity Recognition tasks. |
Copied to clipboard
| Challenge: | a uniform information density hypothesis is used to explain certain linguistic phenomena . a regularizer that encodes the UID hypothesis can be used for language training . |
| Approach: | They propose to augment the canonical MLE objective with a regularizer that encodes UID . they find that regularization consistently improves perplexity in language models . |
| Outcome: | The proposed hypothesis can be operationalized as an inductive bias for language modeling. |
Copied to clipboard
| Challenge: | Conditional models are frequently encountered in practice, but there has not been a rigorous theoretical analysis of NCE in this setting. |
| Approach: | They propose to use a ranking-based and ranking-only method for conditional models to estimate parameter estimates. |
| Outcome: | The proposed method avoids calculation of partition function or derivatives at each training step . it is closely related to negative sampling methods, now widely used in NLP . |
Copied to clipboard
| Challenge: | Variational auto-encoders have been used for text generation but their representation power is limited due to two reasons. |
| Approach: | They advocate sample-based representations of variational distributions for natural language . they further develop an LVM to directly match the aggregated posterior to the prior . |
| Outcome: | The proposed model can be viewed as a natural extension of VAEs with a regularization of maximizing mutual information, mitigating the "posterior collapse" issue. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have been driven not only by advances in neural architectures, but also through hardware and optimization improvements. |
| Approach: | They revisit the neural probabilistic language model (NPLM) of Bengio et al. (2003) which simply concatenates word embeddings within a fixed window and passes the result through a feed-forward network to predict the next word. |
| Outcome: | The proposed model performs better on word-level language model benchmarks than a baseline Transformer with short input contexts but struggles to handle long-term dependencies. |
Copied to clipboard
| Challenge: | Neural network models for many NLP tasks have grown increasingly complex in recent years . authors of recent papers question the necessity of such architectures and find them quite effective . |
| Approach: | They propose to use regularization techniques borrowed from language modeling to improve model accuracy . they find that a simple biLSTM architecture with appropriate regularization yields competitive results . |
| Outcome: | a simple biLSTM model outperforms the state-of-the-art on four benchmark datasets . authors say that improvements are not real, but are attributed to mundane reasons . |
Copied to clipboard
| Challenge: | Existing approaches to train variational autoencoders (VAEs) have been proposed to alleviate the posterior collapse issue in NLP tasks. |
| Approach: | They propose to introduce a mutual information term between the input and its latent variable to regularize the objective of the VAE. |
| Outcome: | The proposed model performs better on three benchmark datasets and is comparable to state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing literature on stereotypical biases in language models is limited . current evaluations focus on measuring bias without considering language modeling ability . |
| Approach: | They propose to measure stereotypical biases in four domains: gender, profession, race, and religion . they compare stereotypical and language modeling ability of popular models like BERT, GPT-2, RoBERTa and XLnet . |
| Outcome: | The proposed model shows strong stereotypical biases in gender, profession, race, and religion domains. |
Copied to clipboard
| Challenge: | Existing multi-view learning models prioritize complementarity while ignoring consensus . EMHA allows for efficient modeling of global dependencies among tokens in parallel . |
| Approach: | They propose an enhanced multi-head self-attention (EMHA) that prioritizes complementarity while ignoring consensus. |
| Outcome: | The proposed method favors consensus among heads by introducing two models . it is superior on a wide range of language tasks with a modest increase in model size . |
Copied to clipboard
| Challenge: | Recent studies show that self-attention patterns in trained models contain a majority of non-linguistic regularities. |
| Approach: | They propose a technique to allow efficient self-supervised learning with bi-directional Transformers by using an auxiliary loss function to guide attention heads to conform to such patterns. |
| Outcome: | The proposed method achieves state-of-the-art in low-resource settings and is agnostic to pre-training objectives. |
Copied to clipboard
| Challenge: | Existing approaches to model long-term dependencies are limited to long texts with thousands of words. |
| Approach: | They propose a look-ahead memory that augments the recurrence memory by attending to the right-side tokens and interpolating with the old memory states to maintain long-term information in the history. |
| Outcome: | Experiments on widely used language modeling benchmarks show that LaMemo outperforms baseline models with recurrence memory. |
Copied to clipboard
| Challenge: | Existing methods require computationally expensive relative position embeddings. |
| Approach: | They propose two methods that decrease input length to improve perplexity and perplexability. |
| Outcome: | The proposed methods speed up training by a factor of 1.65 and reduce memory usage. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a linguistic phenomenon defined as "the alternation of two languages within a single discourse, sentence or constituent." |
| Approach: | They propose an ASR-motivated evaluation setup which is decoupled from an ASL system and the choice of vocabulary . they propose a discriminative training approach which works better than generative language modeling . |
| Outcome: | The proposed evaluation setup is better than generative language modeling, the authors show . the proposed setup is decoupled from an ASR system and the choice of vocabulary . |
Copied to clipboard
| Challenge: | Existing methods to improve neural language models perform poorly on emerging data. |
| Approach: | They propose a lexical-level masking strategy to post-train a neural language model using static data from past years. |
| Outcome: | The proposed method outperforms existing methods on two pre-trained language models, two classification tasks, and four benchmark datasets. |
Copied to clipboard
| Challenge: | Social chatbots evolve rapidly with large pretrained language models. |
| Approach: | They propose effective defense objectives to protect persona leakage from hidden states by a simple neural network. |
| Outcome: | The proposed defense objectives reduce the attack accuracy from 37.6% to 0.5% while preserving language models’ powerful generation ability. |
Copied to clipboard
| Challenge: | Existing research suggests that privacy preservation comes at the price of worsening biases in classification tasks. |
| Approach: | They propose to incorporate privacy preservation and de-biasing techniques into training text generation models to investigate the trade-off between the two dimensions. |
| Outcome: | The proposed model improves on bias detection, privacy attacks, language modeling, and performance on downstream tasks. |
Copied to clipboard
| Challenge: | Recent studies show that pre-trained models are beneficial to Chinese Word Segmentation (CWS). However, these models lack task-specific prior segmentation knowledge. |
| Approach: | They propose a pre-trained Chinese word segmentation model MetaSeg which incorporates meta learning into a multi-criteria pre-training task. |
| Outcome: | Empirical results show that MetaSeg can achieve new state-of-the-art performance on twelve widely-used CWS datasets and significantly improve model performance in low-resource settings. |
Copied to clipboard
| Challenge: | a diversity advanced actor-critical reinforcement learning framework is used to improve NLP generalization and accuracy. |
| Approach: | They introduce Diversity Advanced Actor-Critic reinforcement learning framework to improve NLP generalization and accuracy. |
| Outcome: | The proposed framework outperforms domain adaptation and generalization baselines without using any target domain knowledge. |
Copied to clipboard
| Challenge: | Recent work on structure-aware models have shown promising results on language modeling, but how to incorporate structure knowledge on corpus without syntactic annotations remains an open problem. |
| Approach: | They propose a neural variational language model which enables the sharing of grammar knowledge among different corpora. |
| Outcome: | The proposed model converges significantly faster to lower perplexity on two popular benchmark datasets. |
Copied to clipboard
| Challenge: | Speech and text are two major forms of human language and little effort has been made to model them together. |
| Approach: | They propose to combine speech and text models to create mixed speech-text data by using different tokenizers and automatic metrics to evaluate how well the model mixes speech and texts. |
| Outcome: | The proposed model improves over a speech-only baseline and shows zero-shot cross-modal transferability. |
Copied to clipboard
| Challenge: | State-of-the-art models in natural language processing (NLP) often incorporate sentence encoder functions which generate a sequence of vectors intended to represent the in-context meaning of each word in an input text. |
| Approach: | They conduct the first large-scale systematic study of candidate pretraining tasks, comparing 19 different tasks as alternatives and complements to language modeling. |
| Outcome: | The proposed model can be used to train sentences on language modeling tasks. |
Copied to clipboard
| Challenge: | Structured pruning and knowledge distillation are often not efficient and require a fixed architecture, limiting flexibility. |
| Approach: | They propose a method which integrates knowledge distillation and structured pruning by replacing transformer blocks with smaller, efficient versions during training. |
| Outcome: | The proposed method outperforms L1 pruning and maintains four-fifths of performance on language modeling and commonsense reasoning tasks. |
Copied to clipboard
| Challenge: | Language models perform differently across languages, a new study suggests . morphological typology may explain some of the performance differences, authors say . |
| Approach: | They propose to test morphological alignment of tokenizers, tokenization quality and disparities in dataset sizes and measurement to test this hypothesis. |
| Outcome: | The proposed model shows that fusional languages perform better than fusionative languages . the authors suggest that morphological typology may explain some of the performance differences . |
Copied to clipboard
| Challenge: | Prompt-based learning can tackle zero-shot and few-shot NLP tasks . authors propose a method that makes use of pre-trained language models . |
| Approach: | They propose to map NLP tasks into natural language prompts, which are then filled by pre-trained language models. |
| Outcome: | The proposed method outperforms standard prompt-based methods in few-shot settings. |
Copied to clipboard
| Challenge: | Sentence summarization systems that use latent space to reconstruct the source sentence are unwillingly exploited. |
| Approach: | They propose a method that uses language modeling and semantic similarity metrics to find a high-scoring summary. |
| Outcome: | The proposed method achieves state-of-the-art for unsupervised sentence summarization according to ROUGE scores. |
Copied to clipboard
| Challenge: | Existing retrieval-augmented language models require access to internal representations to enhance performance. |
| Approach: | They introduce a retrieval-augmented language modeling framework that treats the language model as a black box and augments it with a tuneable retrieval model. |
| Outcome: | The proposed framework improves performance on language modeling tasks by 6.3% and 5.1%. |
Copied to clipboard
| Challenge: | Mamba-based SSL models are promising for long-sequence modeling, speech unit extraction, and speech self-supervised learning. |
| Approach: | They propose to use Mamba-based HuBERT models as an alternative to Transformer-based SSL architectures. |
| Outcome: | The proposed models outperform Transformer-based models in language modeling tasks while showing superior performance on streaming ASR. |
Copied to clipboard
| Challenge: | a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say . |
| Approach: | They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary . |
| Outcome: | The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks. |
Copied to clipboard
| Challenge: | Pretrained language models have improved writing assistance functions such as autocomplete, but more complex and controllable writing assistants have yet to be explored. |
| Approach: | They build an intent-guided authoring assistant that follows fine-grained author directives by specifying different writing intents. |
| Outcome: | The proposed system generates output satisfying the author's intent and can be rephrased to their liking. |
Copied to clipboard
| Challenge: | Existing learning-to-route methods suffer from the routing fluctuation issue . with the model scale growing, training speed will go slower and memory requirements are heavy . |
| Approach: | They propose a Mixture-of-Experts technique that can scale up the model size of Transformers with an affordable computational overhead. |
| Outcome: | The proposed method outperforms existing learning-to-route methods on language modeling and multilingual machine translation. |
Copied to clipboard
| Challenge: | a recent study suggests that language models perform poorly across languages. |
| Approach: | They propose a model that fits a paired-sample multiplicative mixed-effects model to obtain language difficulty coefficients from at least-pairwise parallel corpora. |
| Outcome: | The proposed model is able to handle missing data and is aware of inter-sentence variation. |
Copied to clipboard
| Challenge: | Long short term memory units are powerful tools for language modeling, but their performance can be limited by the number of parameters. |
| Approach: | They propose a pyramidal recurrent unit which enables learning representations in high dimensional space with more generalization power and fewer parameters. |
| Outcome: | The proposed model outperforms existing models with different gating mechanisms and transformations on word-level language modeling tasks. |
Copied to clipboard
| Challenge: | Existing language modeling methods rely on large-scale text data to learn the sequential patterns of words. |
| Approach: | They propose to use sememes to represent the implicit semantics behind words for language modeling . they propose to employ sememe-driven language models to fine-grained semem-level semantics . |
| Outcome: | Experiments on language modeling and the downstream application of headline generation show the effectiveness of SDLM. |
Copied to clipboard
| Challenge: | Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed. |
| Approach: | They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic. |
| Outcome: | The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary. |
Copied to clipboard
| Challenge: | Recent work shows that recurrent neural networks can implicitly capture hierarchical information when trained to solve common natural language processing tasks. |
| Approach: | They propose a convolutional sequence-to-sequence model that exploits hierarchical information implicitly. |
| Outcome: | The proposed model is recurrent and non-recurrent, and it can model hierarchical structure implicitly. |
Copied to clipboard
| Challenge: | Existing methods for sentence summarization require a large amount of parallel data for supervision to work. |
| Approach: | They propose an unsupervised method for sentence summarization using only language modeling. |
| Outcome: | The proposed method maintains continuous contextual matching while maintaining output fluency without any paired examples. |
Copied to clipboard
| Challenge: | Existing approaches to attention with bounded-memory control (ABC) have a quadratic complexity in sequence lengths, making it prohibitive for long sequences. |
| Approach: | They propose a new abstraction that bounds memory size to improve efficiency . they propose bounded-memory control, which connects several efficient attention variants . |
| Outcome: | The proposed approach outperforms existing approaches on language modeling, machine translation, and masked language model finetuning. |
Copied to clipboard
| Challenge: | Masked diffusion language models have achieved significant progress in language modeling . however, the systematic analysis and empirical validation of their alignment on general tasks remains underexplored. |
| Approach: | They propose a framework that analyzes the bias and variance of preference optimization loss and gradient based on Direct Preference Optimization. |
| Outcome: | The proposed model outperforms its SFT-only predecessor on general benchmarks . it consistently outperformed other strong language models and ARMs on general tasks . |
Copied to clipboard
| Challenge: | Existing methods to train language models have limitations in interpretability . a Word-Context-Coupled Space (W2CSpace) is proposed to improve the performance of pre-trained models . |
| Approach: | They propose a Word-Context-Coupled Space to replace pre-trained models with interpretable statistical logic. |
| Outcome: | The proposed language model can achieve better performance and highly credible interpretability compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing approaches to NLP are sparsifying attention patterns or approximating the attention computation with kernel methods. |
| Approach: | They propose a method for dynamic contextual compression for decoder-only LMs. |
| Outcome: | The proposed method reduces the cost of self-attention to a fraction of typical time and space. |
Copied to clipboard
| Challenge: | Language modeling is a core task in natural language processing. |
| Approach: | They propose to characterize leakage onto the set of infinite sequences by a measure-theoretic approach. |
| Outcome: | The proposed language model families are tight, meaning they will not leak . the proposed language models are based on the 'sequence leakage' hypothesis . |
Copied to clipboard
| Challenge: | Recent work on latent tree learning attempts to develop models with parse-valued latent variables and train them on non-parsing tasks. |
| Approach: | They propose a model with parse-valued latent variables and a strong latent tree learning result on constituency parsing. |
| Outcome: | The proposed model outperforms all baselines and performs competitively with symbolic grammar induction systems. |
Copied to clipboard
| Challenge: | Variational autoencoders (VAEs) are a popular family of generative models with wide applicability. |
| Approach: | They propose to modify a deterministic model designed for images to avoid posterior collapse by controlling the entropy of the aggregate posterior to make it Gaussian. |
| Outcome: | The proposed models outperform a broad range of VAE models on text generation and downstream tasks from representations while avoiding reparametrization steps. |
Copied to clipboard
| Challenge: | Large language models exhibit reasonable multilingual abilities, despite predominantly English-centric pretraining. |
| Approach: | They propose a framework that establishes multilingual alignment prior to language model pretraining and preserves this alignment using a code-switching strategy during pretraining. |
| Outcome: | Experiments in a synthetic English to English-Clone setting show that PreAlign outperforms standard multilingual joint training in language modeling, zero-shot cross-lingual transfer, and cross-linguistic knowledge application. |
Copied to clipboard
| Challenge: | Existing language modeling datasets contain near-duplicate examples and long repetitive substrings. |
| Approach: | They develop tools that allow us to deduplicate existing language modeling datasets . they found that over 1% of the unprompted output of language models is copied verbatim . |
| Outcome: | The proposed tools reduce train-test overlap, which affects over 4% of validation sets, and improve model accuracy. |
Copied to clipboard
| Challenge: | a number of RNNs update their state as the input sequence is processed . second-order RNN architectures show promising performance in language modeling . |
| Approach: | They propose a second-order RNN architecture that generalizes existing ones . they use a Penn Treebank dataset to analyze how their different components affect performance . |
| Outcome: | The proposed architecture generalizes existing RNNs on a Penn Treebank dataset . it shows that removing the first-order terms does not hinder performance . |
Copied to clipboard
| Challenge: | Evidence has shown that multi-head attentive neural architectures are overparameterized. |
| Approach: | They propose a multi-head attentive neural architecture that “reallocates” attention heads to different inputs. |
| Outcome: | The proposed model outperforms baselines on machine translation and language modeling tasks. |
Copied to clipboard
| Challenge: | Existing general purpose components for learning differentiable windows are hard to optimize. |
| Approach: | They propose a new neural module and general purpose component for dynamic window selection that can enable more focused attentions over the input regions. |
| Outcome: | The proposed approach improves on a myriad of NLP tasks including machine translation, sentiment analysis, subject-verb agreement and language modeling. |
Copied to clipboard
| Challenge: | In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families. |
| Approach: | They present a set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. |
| Outcome: | The Bloom Library datasets cover 363 languages across 32 language families. |
Copied to clipboard
| Challenge: | incorporating syntactic structure into language models has been a challenge since the 1990s. |
| Approach: | They propose to use syntactic information to integrate syntastic structure into neural language models by providing ground truth parse trees as additional training signals. |
| Outcome: | The proposed model achieves lower perplexity and better quality when ground truth parse trees are provided as training signals. |
Copied to clipboard
| Challenge: | Large-scale multitask benchmarks have driven rapid progress in language modeling, yet most emphasize low-resource languages like English. |
| Approach: | They propose a benchmark for massive multitask language understanding in Bengali . they use a dataset that preserves mathematical content via MathML and a subset of questions most frequently missed by top systems to stress difficult cases. |
| Outcome: | The proposed benchmark covers 24 model variants across 11 LLM families. |
Copied to clipboard
| Challenge: | Using a key-value cache, memory consumption is a bottleneck for high-throughput language models. |
| Approach: | They propose a method that only computes and caches the KVs of a small number of layers, thus saving memory consumption and improving inference throughput. |
| Outcome: | The proposed method achieves higher throughput and competitive performance than standard transformers and is orthogonal to existing transformer memory-saving techniques. |
Copied to clipboard
| Challenge: | a subset of words belonging to specific psycholinguistic categories vary more in their representations across users . combining generic and personalized word embeddings yields the best performance . |
| Approach: | They propose personalized word embeddings and compare their performance to generic ones . they show that personalized word representations can be leveraged for improved performance . |
| Outcome: | The proposed model can be used for authorship attribution. |
Copied to clipboard
| Challenge: | Reinforcement learning from human feedback (RLHF) is the primary method for aligning large language models with human preferences. |
| Approach: | They propose to train an Absolute-Rating Multi-Objective Reward Model with multi-dimensional absolute-rating data. |
| Outcome: | The proposed model outperforms the LLM-as-a-judge method on RewardBench . it achieves state-of-the-art performance on the benchmark . |
Copied to clipboard
| Challenge: | Neural architecture search (NAS) uses weight-sharing supernets to generate diverse subnetworks without retraining. |
| Approach: | They propose a weight-sharing supernet that leverages mixture-of-experts to enhance supernet model expressiveness with minimal training overhead. |
| Outcome: | The proposed method achieves state-of-the-art (SoTA) performance in NAS for fast machine translation models, surpassing NAS-BERT and AutoDistil across various model sizes. |
Copied to clipboard
| Challenge: | Recent approaches to language model alignment assume homogeneous human preferences, but actual human preferences vary widely and are hard to satisfy with a single language model. |
| Approach: | They propose an RL-free extension of Direct Preference Optimization (DPO) that folds language modeling directly into reward modeling and trains language models as collective reward models that combine all objectives with specific weights. |
| Outcome: | The proposed method matches or outperforms existing methods in safety alignment and long-form question answering. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have led to computationally intensive and resource-demanding state-of-the-art models. |
| Approach: | They investigate the impact of pre-training data volume on compact language models . they use a French question answering task to train models with as little as 100 MB of text . |
| Outcome: | The results show that pre-training data volume can improve models with as little as 100 MB of text . the results suggest that the model performance is poorer with less data than with larger datasets . |
Copied to clipboard
| Challenge: | Recent advances in NLP demonstrate the effectiveness of training large-scale language models and transferring them to downstream tasks. |
| Approach: | They conduct an extensive study of the transferability between 33 NLP tasks across three broad classes of problems. |
| Outcome: | The proposed model can improve performance even with low-data source tasks that differ substantially from the target task. |
Copied to clipboard
| Challenge: | State-of-the-art reading comprehension models do not have general linguistic intelligence . accuracy of out-domain datasets is affected by the distribution of data . |
| Approach: | They propose to use supervised RC training data in the source domain and unlabeled passages in the target domain to adapt models. |
| Outcome: | The proposed model outperforms the model without domain adaptation with five datasets in different domains. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have enabled new conversational systems. |
| Approach: | They propose to use a dataset of indirect referring expressions to solve the problem of reference resolution when people use natural expressions . they propose to model the problem using 42K indirect referred expressions across three domains and a public dataset of entity pairs and utterances. |
| Outcome: | The proposed models achieve 82%-87% accuracy in realistic settings, while reasonable invites further advances. |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have achieved competitive performance on a range of NLP tasks. |
| Approach: | They propose to learn distributional invariance across source domains via alignment regularization loss functions to improve domain generalization by prompting. |
| Outcome: | Experiments on sentiment analysis and natural language inference show the effectiveness of the proposed method and achieve state-of-the-art results. |
Copied to clipboard
| Challenge: | Large-scale generative Pre-trained Language Models (PLMs) are limited in their deployment in real-world applications. |
| Approach: | They propose to prune the feed-forward networks of generative pre-trained language models to smaller widths without designing extra operators. |
| Outcome: | The proposed method achieves 1.51x/6.96x inference speedup on GPU/CPU with 67% size reduction. |
Copied to clipboard
| Challenge: | XMoE leverages small experts and a threshold-based router to selectively engage only essential parameters. |
| Approach: | They propose a novel MoE that leverages small experts to selectively engage only essential parameters. |
| Outcome: | The proposed model can reduce computation load at MoE layers by over 50% without sacrificing performance. |
Copied to clipboard
| Challenge: | Continual event detection (CED) is a challenging task due to catastrophic forgetting, where learning new tasks hampers performance on previous ones. |
| Approach: | They propose a method that leverages optimal transport principles to align the optimization of a classification module with the intrinsic nature of each class, as defined by their pre-trained language modeling. |
| Outcome: | The proposed method outperforms state-of-the-art methods on MAVEN and ACE datasets and is a pioneering solution in continual event detection. |
Copied to clipboard
| Challenge: | Existing work on the authorship of the Homeric poems has only been done at the level of lengthier excerpts, but not individual verses, at which most suspected interpolations occur. |
| Approach: | They present a corpus of Homeric verses with a score quantifying linguistic unexpectedness based on Perplexity. |
| Outcome: | The proposed corpus of Homeric verses is complemented with a score quantifying linguistic unexpectedness based on Perplexity. |
Copied to clipboard
| Challenge: | Currently, attention-based models face computational hurdles in processing long sequences due to its quadratic complexity. |
| Approach: | They propose a conformer whose encoder self-attentions are replaced with Hyena for speech processing . they propose 'confhyena' model that reduces training time by 27% at minimal cost . |
| Outcome: | The proposed model reduces training time by 27% at the cost of minimal quality degradation. |
Copied to clipboard
| Challenge: | Quantization studies have focused on instruction-tuned LLMs, leaving their performance on other benchmarks unclear. |
| Approach: | They propose a framework to evaluate quantized large language models using four dimensions . they propose to reduce the bits needed for model weights or activations with minimal performance loss . |
| Outcome: | The proposed framework can retain comparable performance to non-quantized LLMs on most benchmarks. |
Copied to clipboard
| Challenge: | Existing studies have argued that function words aid learning abstract grammatical knowledge from linear input. |
| Approach: | They examine the statistical distribution of function words and their properties . they show that function words are reliable, diverse, and informative . |
| Outcome: | The results show that function words preserve high frequency, reliable syntactic association, phrase-boundary alignment and are informative to structural dependency. |
Copied to clipboard
| Challenge: | Existing LLMs are constrained by their pre-trained context lengths, leading to performance issues . elucidating this limitation, we propose a training-free solution to the context length limitation in LLM applications . |
| Approach: | They propose a method that integrates hierarchical rotary position embedding into LLMs without extra training costs. |
| Outcome: | The proposed method improves performance on language modeling and long code completion tasks. |
Copied to clipboard
| Challenge: | Prompt-based learning is susceptible to intrinsic bias present in pre-trained language models (LMs), leading to sub-optimal performance in prompt-based zero/few-shot settings. |
| Approach: | They propose a null-input prompting method to calibrate intrinsic bias encoded in pre-trained language models (LMs) they leverage a diverse set of auto-selected null meaning inputs generated from GPT-4 to probe intrinsic bias. |
| Outcome: | The proposed method significantly improves zero/few-shot learning performance of LMs for both in-context learning and prompt-based fine-tuning (on average 9% and 2%, respectively). |
Copied to clipboard
| Challenge: | Toeplitz Neural Networks outperform commonly used Transformer-based models while benefiting from log-linear space-time complexities. |
| Approach: | They propose to convert TNNs to SSMs during inference to combine strengths of TNN and SSM approaches. |
| Outcome: | The proposed method outperforms most Transformer-based models while retaining the advantage of constant inference complexity. |
Copied to clipboard
| Challenge: | Existing approximations of dot-product attention ignore the value vectors . a value-aware objective outperforms an optimal approximate that ignores values . |
| Approach: | They propose an approximation of a value-aware objective that substantially outperforms an optimal approximate that ignores values. |
| Outcome: | The proposed value-aware objective outperforms an optimal approximation that ignores values in the context of language modeling. |
Copied to clipboard
| Challenge: | Existing language modeling tools for automatic speech recognition (ASR) are difficult to personalize. |
| Approach: | They propose a domain-distributed Span-Aggregated K-nearest N-gram retrieval augmentation to improve language modeling for automatic speech recognition (ASR) personalization. |
| Outcome: | The proposed model outperforms baselines on Wikitext-103, UserLibri, and ASAP datasets with a 10-16% improvement in perplexity and a 5-8% reduction in word error rates. |
Copied to clipboard
| Challenge: | In natural language processing, the representation of text plays a crucial role in various tasks such as language modeling, sentiment analysis, and machine translation. |
| Approach: | They propose a method to represent English text with only consonants that is more discriminative than vowels and a technique to retrieve vowel information from it. |
| Outcome: | The proposed representation significantly reduces the overall memory and compute footprint required for storing and processing textual data. |
Copied to clipboard
| Challenge: | Memorization is a fundamental ability of Transformer-based Large Language Models, achieved through learning. |
| Approach: | They propose an architecture that explicitly memorizes sequences of tokens in layered associative memories. |
| Outcome: | The proposed architecture shows that memorization is a fundamental ability of large language models, achieved through learning. |
Copied to clipboard
| Challenge: | Quantization has proven to be effective after pre-training and during fine-tuning, but its effects on pre-trainer performance have remained unexplored. |
| Approach: | They propose a linear quantization strategy to be applied during the pre-training of Transformers to improve model efficiency and stability. |
| Outcome: | The proposed method improves model efficiency, stability, and performance while maintaining language modeling ability. |
Copied to clipboard
| Challenge: | Sequence models produce accurate predictions, but their decision making processes are hard to explain. |
| Approach: | They propose an efficient algorithm to approximate sequential objective by identifying the most faithful rationales. |
| Outcome: | The proposed algorithm is best at optimizing the sequential objective and provides the most faithful rationales. |
Copied to clipboard
| Challenge: | Recent work on large language models has made this hypothesis popular . but, word order is not important enough to make sentence structure relevant . |
| Approach: | They propose an efficient procedure that finds word order having highest likelihood under a fixed language model. |
| Outcome: | The proposed procedure can be used to find the ordering of a bag of words having the highest likelihood under a fixed language model. |
Copied to clipboard
| Challenge: | Existing Transformers can only deal with the in-distribution size of inputs. |
| Approach: | They propose a relative position embedding to explicitly maximize attention resolution . they also use blockwise causal attention during inference for better resolution a . |
| Outcome: | The proposed model achieves strong performance in interpolation and extrapolation settings. |
Copied to clipboard
| Challenge: | Multilingual models have been released, but many of the world's languages are not covered. |
| Approach: | They propose a method that initializes the embedding matrix for a new tokenizer based on information in the source model's embeddable matrix. |
| Outcome: | The proposed method outperforms random initialization and previous work on language modeling and on a range of downstream tasks (NLI, QA, and NER). |
Copied to clipboard
| Challenge: | Current NLP models produce unreliable or catastrophic predictions when training and test distributions differ . current models tend to produce unreliability or even catastrophic predictions that hurt user trust. |
| Approach: | They categorize examples as exhibiting a background shift or semantic shift and use calibration and density estimation methods to detect OOD examples. |
| Outcome: | The proposed methods beat calibration methods in background shift settings and perform worse in semantic shift settings. |
Copied to clipboard
| Challenge: | Transformer-based language models create hidden representations of inputs at every layer, but only use final-layer representations for prediction. |
| Approach: | They propose a method for casting hidden representations as final representations, bypassing transformer computation in-between. |
| Outcome: | The proposed method produces more accurate predictions from hidden layers across various model scales, architectures, and data distributions. |
Copied to clipboard
| Challenge: | Existing work on multilingual summarization and cross-lingual summmarization has been limited due to their different definitions. |
| Approach: | They propose to unify MLS and CLS into a more general setting, i.e. many-to-many summarization. |
| Outcome: | The proposed model outperforms the state-of-the-art models in the zero-shot directions. |
Copied to clipboard
| Challenge: | Existing decoding strategies for pre-trained MDLMs rely on token-level uncertainty criteria, while largely overlooking sequence-level information and inter-token dependencies. |
| Approach: | They propose a training-free decoding strategy that leverages inter-token dependencies to inform token updates during generation. |
| Outcome: | Empirical results show that the proposed approach consistently achieves superior performance on both code generation and mathematical reasoning tasks. |
Copied to clipboard
| Challenge: | Contemporary advances in NLP are built on the representational power of latent embedding spaces learned by self-supervised language models (LMs). |
| Approach: | They use a new information theoretic probing suite to analyze representational subspaces in language models. |
| Outcome: | The proposed approach compared performance of nine tasks across 2M pre-training steps and five seeds. |
Copied to clipboard
| Challenge: | Subword tokenization methods are often used to project subwords onto triplets . a typical tokenizer consists of 10 000s of subword mapped onto a single index . |
| Approach: | They propose a subword tokenization method that factorizes subwords onto triplets using a VQ-VAE model. |
| Outcome: | The proposed tokenization method is more appropriate and robust for morphological tasks than the commonly used byte-pair encoding (BPE) tokenization algorithm. |
Copied to clipboard
| Challenge: | Prior work on compression prioritizes preserving perplexity, which is analogous to training loss. |
| Approach: | They examine the impact of model compression along four dimensions: degeneration harm, representational harm, dialect bias, and language modeling and downstream task performance. |
| Outcome: | The proposed compression methods can lead to unexpected consequences, the authors show . quantization preserves bias while pruning degrades quickly. |
Copied to clipboard
| Challenge: | Large and sparse feed-forward layers (S-FFN) have proven effective in scaling up the model size for pretraining large language models. |
| Approach: | They compare S-FFN architectures for language modeling and compare their performance and efficiency . they found a simpler selection method that selects blocks through their mean aggregated hidden states . |
| Outcome: | The proposed model size and selection method achieve lower perplexity in language model pretraining compared to existing MoE architectures. |
Copied to clipboard
| Challenge: | Experimental results show VocalNet outperforms existing open-source speech LLMs despite limited training data. |
| Approach: | They propose a scalable and model-agnostic training framework and a novel multi-token prediction paradigm for speech generation. |
| Outcome: | The proposed model outperforms open-source speech LLMs while outperforming existing open-sourced models. |
Copied to clipboard
| Challenge: | In-context learning (ICL) is an emerging capability of large autoregressive language models where a few demonstrations are appended to the input to enhance the model’s understanding of downstream NLP tasks without directly adjusting the model parameters. |
| Approach: | They propose a method where a few demonstrations are appended to the input to enhance the model's understanding of downstream NLP tasks without directly adjusting the model parameters. |
| Outcome: | The proposed method significantly improves the input-label mapping in ICL demonstrations. |
Copied to clipboard
| Challenge: | Existing sparsification methods like pruning can lose model knowledge through parameter removal. |
| Approach: | They propose a novel approach that achieves sparsification by partitioning pre-trained FFN layers into computational blocks. |
| Outcome: | The proposed approach achieves superior performance across language modeling and downstream tasks under equivalent computational constraints. |
Copied to clipboard
| Challenge: | Current research on bias in language models focuses on data quality, not temporal influences of data. |
| Approach: | They propose a methodology to interpret the interaction between training data and model architecture in bias propagation during language modeling. |
| Outcome: | The proposed method analyzes the interaction between training data and model architecture in bias propagation during language modeling. |
Copied to clipboard
| Challenge: | OpenAI’s o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts. |
| Approach: | They curate a small dataset s1K with 1,000 reasoning questions based on three criteria we validate through ablations: difficulty, diversity, and quality. |
| Outcome: | The proposed model exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). |
Copied to clipboard
| Challenge: | kNN-LM, REALM, DPR + FiD, Contriever + ATLAS, and Contriver + Flan-T5 are popular retriever-augmented language models for a variety of tasks. |
| Approach: | They evaluate the strengths and weaknesses of kNN-LM, REALM, DPR + FiD, Contriever + ATLAS, and Contriver + Flan-T5 in reasoning over retrieved statements across different tasks. |
| Outcome: | The proposed models do not exhibit strong reasoning even when provided with only the required statements. |
Copied to clipboard
| Challenge: | Recent efforts to "personalize" large language models by assigning them specific personas are limited by current knowledge of how well they perform. |
| Approach: | They use a style embedding model to analyze writing styles of persona-assigned LLMs . they find significant style differences between personas using Kullback-Leibler divergence . |
| Outcome: | The proposed model shows significant differences in writing styles among personas across socio-demographic groups. |
Copied to clipboard
| Challenge: | Existing approaches to training language models for each jurisdiction fail to leverage common legal principles beneficial for low-resource settings or risk negative interference from conflicting jurisdictional interpretations. |
| Approach: | They propose a parameter-efficient framework that derives hierarchical relationships across jurisdictions and progressively inserts adapter modules across model layers based on jurisdictional similarity. |
| Outcome: | The proposed framework outperforms fully shared and jurisdiction-specific models on two legal language modeling benchmarks. |
Copied to clipboard
| Challenge: | Autoregressive (AR) language models are a dominant paradigm in the field of parallelism and non-causal modeling. |
| Approach: | They propose a blockwise discrete diffusion model that preserves AR-compatible serving while enabling parallel intra-block generation. |
| Outcome: | The proposed model achieves theoretical speedups over 5 and wall-clock speedup of 2.3 on H200 GPUs in latency-critical regimes. |
Copied to clipboard
| Challenge: | Conventional statistical tokenizers often disrupt constituent boundaries within words, thereby corrupting semantic information. |
| Approach: | They propose a method that uses morphological structure guidance to induce character-level structures of words by training a deep model. |
| Outcome: | Empirical results show that the proposed method retains complete morphemes and outperforms existing methods on morphological segmentation and language modeling tasks. |
Copied to clipboard
| Challenge: | Existing defenses for large language models do not account for the sequential nature of text data. |
| Approach: | They propose a lightweight yet effective empirical privacy defense that leverages token-specific characteristics to protect training data of large language models. |
| Outcome: | The proposed approach provides strong protection against membership inference attacks and improves language modeling performance by 10% across different LLM architectures and datasets compared to baselines. |
Copied to clipboard
| Challenge: | Large-scale language models (LLMs) have shown remarkable performance across a wide array of tasks. |
| Approach: | They propose an architecture that preserves parameter efficiency of tied models without sacrificing representational benefits of untied embeddings. |
| Outcome: | The proposed architecture achieves a 31.72% improvement in linguistic knowledge acquisition over the baseline model. |
Copied to clipboard
| Challenge: | Existing methods to detect text spans that refer to entities are often conflated with entity typing in a single joint task. |
| Approach: | They propose a lightweight model that probes mention detection capabilities from early LLM layers. |
| Outcome: | The proposed model achieves 93% recall zero-shot with 90% precision under human-calibrated LLM-judge protocol . |
Copied to clipboard
| Challenge: | Data selection techniques have shown empirical benefits in reducing the number of gradient steps to train neural models. |
| Approach: | They propose to modify an existing data selection technique to adapt it to the sequence losses typical in language modeling. |
| Outcome: | The proposed technique reduces the number of steps required to train neural models by 4.3% and improves generalization ability on out of domain datasets. |
Copied to clipboard
| Challenge: | MoE-based LLMs are not explicitly supervised to select suitable experts. |
| Approach: | They propose Exploration-Driven Reinforcement Learning (ERL) which explicitly optimizes the router by exploration of alternative routing paths. |
| Outcome: | The proposed method improves summarization (SAMSum, XSUM, question answering, and language modeling), and raises routing quality, delivering 8.9 higher MRR than baselines over 100 perturbed routing paths. |
Copied to clipboard
| Challenge: | Existing static vocabulary pruning designs that reduce memory usage suffer from rigid, one-size-fits-all designs that cause information loss during the prefill stage and lack flexibility. |
| Approach: | They propose a decoupled dynamic vocabulary selection framework that addresses memory constraints through offloading embedding and implements a hybrid static-dynamic vocabulary selection strategy for LM Head. |
| Outcome: | The proposed framework reduces memory usage by 99% with minimal or no degradation in performance. |
Copied to clipboard
| Challenge: | Prior research on linguistic mechanisms of large language models is limited by coarse granularity, limited analysis scale, and narrow focus. |
| Approach: | They propose a framework for analyzing the linguistic mechanisms of large language models based on Sparse Auto-Encoders. |
| Outcome: | The proposed framework extracts Chinese and English linguistic features across four dimensions . it uncovers intrinsic representations of linguistic knowledge in LLMs and can control outputs . |
Copied to clipboard
| Challenge: | Recent work focuses on syntactic tree structures of languages, in particular constituency tree structures. |
| Approach: | They propose a Graph-Infused Layers Transformer Language Model which leverages dependency graphs to augment Transformer language models. |
| Outcome: | The proposed model achieves better syntactic generalization while maintaining competitive perplexity compared with baseline models. |
Copied to clipboard
| Challenge: | a recent study shows that subword tokenization improves performance of neural language models. |
| Approach: | They propose a linguistically grounded approach to train a tokenizer on morphologically segmented data. |
| Outcome: | The proposed tokenizer improves on a Spanish language model with morphological information. |
Copied to clipboard
| Challenge: | Recent advances in language modeling have shown promising results when applied to time series data. |
| Approach: | They propose a method to fine-tune large language models for time series classification tasks using text embedding models and a simple classification head. |
| Outcome: | The proposed model outperforms the current SOTA model on a time series classification benchmark and uses only 14.5% of the trainable parameters. |
Copied to clipboard
| Challenge: | Existing methods for implementing large language models are limited by high computational and memory requirements. |
| Approach: | They propose a lightweight binarization framework that achieves effective W(1+1)A4 quantization through a novel three-stage quantization strategy. |
| Outcome: | The proposed framework surpasses state-of-the-art methods on W2A4 quantization settings across languages. |
Copied to clipboard
| Challenge: | Structured pruning is a practical approach to deploying large language models (LLMs) but it fails to capitalize on modest task-specific calibration signals, causing limited downstream gains. |
| Approach: | They propose a method that removes attention heads and MLP channels using loss-based important scores . they use perplexity for language modeling and a margin-based objective for decision-style tasks . |
| Outcome: | The proposed method lowers perplexity and improves accuracy at higher sparsity . it also stabilizes accuracy and mitigates perxity collapse without fine-tuning . |
Copied to clipboard
| Challenge: | Existing studies show that global and local attention are expressively complementary. |
| Approach: | They propose to restrict global attention to a fixed-size window of preceding tokens . they also propose to add local attention to local-only transformers to increase model quality . |
| Outcome: | The proposed model outperforms the global–local transformers on formal language recognition and natural language modeling. |